IndexPDF journal
Can AI Evaluate a Book Index?
A book-scale Oxford study shows how source-grounded AI evaluation can test a finished index and reach an explicit publication-readiness decision.
AI can evaluate a book index—if its judgments are tied to the source and its publication-readiness rules are explicit.
That distinction matters because evaluating an index is not the same as checking whether its page numbers exist or whether it resembles another index. A subject index makes thousands of claims about a book: that a concept matters, that a heading represents it accurately, that each locator leads to useful treatment, and that the resulting structure helps a reader navigate.
A book-scale test
We applied a source-grounded, AI-assisted evaluation to the supplied published index for William Doyle’s 2002 second edition of The Oxford History of the French Revolution. The public Oxford study covers 425 supplied body-text pages, 1,904 index records, 5,338 atomic locator claims, 16 cross-references, and a source benchmark prepared before the candidate was judged.
The evaluation examined the index in both directions. Candidate-to-source review tested whether the entries and locators were supported. Source-to-candidate review asked whether important subjects and expected treatments could actually be found. A whole-index pass then examined hierarchy, density, cross-references, mechanics, and reader navigation.
That apparent tension is the point. The score is a weighted summary, not “percent correct.” The failed gates identify defects that the method does not allow stronger results elsewhere to cancel:
- One heading misstated a consequential relationship.
- The same compound heading was unsupported as written.
- One cross-reference remained unresolved.
- A localized clutter pattern triggered the standard rule.
The displayed view also corrects eighteen confirmed digit-for-accent substitutions in the published-index PDF distributed by IndexerLabs.¹
A locator also has to do more than contain the right word. As the Garonne example shows, every page reference makes its own promise to the reader. The Oxford audit tests that promise against the complete heading path, not an isolated term.
Two kinds of evidence—and a decision
- Mechanical validation. Software can map pages, expand ranges, validate schemas, account for denominators, resolve identifiers, and reproduce declared arithmetic.
- Source-grounded evaluation. Evidence-grounded review assesses whether passages provide useful treatment, whether headings preserve meaning and stance, and whether important access is missing.
- Publication-readiness decision. Noncompensatory gates determine whether consequential defects block publication, regardless of the overall score.
The Subject Index Evaluation Standard keeps these layers separate. The score is calculated from frozen ledgers; diagnostic item grades do not add up to it, and publication gates remain outside the score.
Why evaluation can be harder than generation
Generation proposes one possible navigation system. Evaluation must verify every proposed claim and search the book for important access that was never proposed. It has to build and review a source-first benchmark, expand every range into individual page assignments, audit the candidate in both directions, inspect the index as a whole, preserve uncertainty, and validate the calculation.
That structure explains why evaluation can require more computation than generation, especially when several candidates share the same benchmark. The Oxford work did not record comparable token use, hardware, cost, or human minutes for generation and evaluation, however, so it does not establish an empirical cost ratio.
So, can AI evaluate a book index?
Yes. In the Oxford study, AI agents reviewed source evidence at book scale, applied a declared rubric, surfaced contradictions, and produced complete audit ledgers. Deterministic software then validated the ledgers and calculated the score and gate outcomes.
The conclusion remains scoped. This is one English-language scholarly history and one candidate, not a validation corpus. The method has not yet been calibrated across genres, model judgments can vary, and live-reader testing measures something the evaluation does not.
There is also a direct competing interest: I created the methodology and develop IndexPDF. The benchmark was prepared in candidate-unseen contexts, but the study is not institutionally independent. Its defense is therefore transparency: the long-form methodology, evaluation framework, frozen benchmark, public result artifacts, and item-level findings are available for scrutiny.