Skip to main content
IndexPDF

IndexPDF journal

The Impossible Third Index: How AI Changed IndexPDF

The story behind IndexPDF—and how its long-running editorial workflow turns final PDF proofs into substantially finished, evidence-grounded book indexes.

Published Updated
Old Testament Use of Old Testament beside pages from its indexes and a laptop displaying IndexPDF.
The assignment that became Djinndex: a 1,104-page reference work with nearly 300 pages of subject, author, and Scripture indexes.

Djinndex began with three wishes: faster subject, author, and Scripture indexes. The third seemed beyond the reach of software. It no longer does.

On December 18, 2020, Zondervan sent me the final draft of the largest book I had ever agreed to index: Gary Edward Schnittjer’s Old Testament Use of Old Testament, a 1,104-page reference work about the Hebrew Bible’s reuse of its own texts. Its pages are dense with biblical citations, scholarly references, and tables. Its subject, author, and Scripture indexes would together occupy nearly 300 pages.

By then I had spent more than a year building software to help me finish it. My son was due in March. The index—and the software—were not yet done. He arrived first. My son was born on March 5, 2021. I delivered the indexes about a week later.

That timeline is less miraculous than the version I sometimes remembered, in which the final book and the finished program arrived together and I produced everything in a weekend. My email history tells a better story. I had early versions of the manuscript, so I could develop the software against the actual book as it evolved. When the final draft arrived, I spent nearly three more months finishing the system, completing the largely hand-built subject index, and preparing the indexes for delivery. I later sent several rounds of corrections as I found things I wanted to improve.

Djinndex did not index an enormous book with the wave of a wand. It gave one person enough leverage to finish a project that might otherwise have taken months. It also taught me exactly where software stopped being useful.

Today, that project is IndexPDF. Modern language models can now do the one thing Djinndex could not: take part in constructing a real subject index. IndexPDF prepares the book, builds and reconciles index material in stages, edits the subject index as a whole, grounds its locators in the source, and gives a human the evidence needed for final approval.

The impossible third index has become possible. The harder question is what we should do with it.

Three wishes

Zondervan had recruited me for the Schnittjer project after I indexed the revised edition of John Goldingay’s Daniel in the Word Biblical Commentary series. While working on Daniel, I wrote shell scripts to find author names and Scripture references automatically.

For Schnittjer’s book, those experiments became Djinndex—a portmanteau of djinn and index. The name referred to three wishes: faster and better subject, author, and Scripture indexing.

Two of the wishes could be expressed as rules. I wrote parsers for the complicated ways Scripture references are written: abbreviations, chapter and verse ranges, discontinuous citations, and combinations of references. I wrote bibliography parsers that turned entries into structured data so the software could recognize an author in the many forms that might appear in the body or notes.

The other major advance was less glamorous and just as important: preparing the book’s text. A PDF does not necessarily contain sentences in the order a person sees them. Headers, footers, footnotes, multiple columns, and tables can turn an apparently orderly page into scrambled input. Schnittjer’s book contained complex, information-dense tables. Extracting them without destroying their meaning became part of the indexing problem.

This is still one of IndexPDF’s strengths. Before any parser or language model can interpret a book, the software has to recover the book that is actually on the page.

Three-stage product history showing a command-line Scripture parser, the blue-and-gold Djinndex application, and the current IndexPDF workflow from automated editorial processing through review and Markdown export.

Before IndexPDF came Djinndex, a web application built on custom algorithms and parsers. Before Djinndex came command-line author and Scripture extractors driven by regular expressions of nearly illegible complexity. The interfaces evolved, but the deeper transformation happened in the technology beneath them.

Before IndexPDF came Djinndex, a web application built on custom algorithms and parsers. Before Djinndex came command-line author and Scripture extractors driven by regular expressions of nearly illegible complexity. The interfaces evolved, but the deeper transformation happened in the technology beneath them.

Djinndex became unusually capable on the book for which I built it. That was both its strength and its weakness. I tried to make the system general, but my first priority was finishing this particular assignment. It was a working tool, not yet a product. And my third wish remained stubbornly out of reach.

The index that was not in the text

A Scripture index can be extracted because the references are present, even when they are difficult to recognize. An author index works much the same way. Subject indexes are different.

A subject index is not hidden inside a book waiting for the right parser. A passage may be about exile without using the word exile. A term may appear on fifty pages but deserve only three locators. Two passages may use different language for the same idea; two similar phrases may name importantly different ideas. The indexer must decide what matters, how it fits together, and what a future reader might call it.

A concordance tells you where a word appears. An index tells you where an idea matters.

I tried increasingly sophisticated ways to bridge that gap. At one point I built a vector database containing the roughly 1.8 million headings in FAST, the subject vocabulary derived from the Library of Congress Subject Headings. The system could associate passages with headings remarkably well; but the headings were still not enough.

The problem was not that 1.8 million possibilities constituted a small vocabulary. It was that the possible subjects of books are effectively unbounded. Back-of-book indexes routinely require distinctions more specialized than a vocabulary designed to organize entire libraries. A universal list could tell me what a passage resembled. It could not decide what this particular book was trying to say.

My third wish required interpretation, and I came to believe that software could not grant it.

The detour that was not a detour

By the time I delivered the indexes, I had spent seven years in graduate school and become a PhD candidate in Old Testament at Fuller Theological Seminary. I also had substantial student debt and a newborn son. During the year I built Djinndex, software engineering had steadily displaced my doctoral work. After my son was born, it became hard to regard that displacement as a hobby.

I applied for software jobs and was hired by Box. For the next four and a half years, I worked on web applications, testing, and product interfaces used at enormous scale.

I thought of the work as preparation. Djinndex had shown me that the idea was valuable, but also how far a tool made for one person and one book was from a product others could trust. A real indexing system would have to survive irregular documents, expose its evidence, make thousands of small decisions manageable, and help users find errors the system itself could not recognize.

By the time I was ready to return, software engineering was undergoing its own upheaval. Technological change is usually described from a distance. A new capability appears, a productivity curve rises, and “workers” move from one category to another. From inside a life, those categories are years: skills acquired, debts incurred, identities assembled, and plans made for people one loves.

I had already watched social change reshape the academic world in which I trained. Then AI began reshaping software engineering. Now it was bringing me back to indexing—but to a form of indexing very different from the one I had left.

That history makes me reluctant either to minimize what AI can do or to celebrate disruption carelessly. The capability is real, but so are the consequences for us “workers.”

Then the threshold moved

The result that changed my mind did not come from IndexPDF. I encountered IndexerLabs and spent an entire weekend testing and analyzing its subject-indexing capabilities, increasingly astonished by what I found. The indexes had weaknesses, but they approached professional quality. They bore the shape of editorial judgment: selecting, consolidating, and organizing ideas in ways I had believed software could not.

At first, I assumed that matching IndexerLabs would require specialized training on a large collection of professional indexes. That raised an immediate question: published indexes embody intellectual judgment, and I was uncertain what rights I would have to use them as training material. But before training a model, a rule of thumb is to first establish that training is necessary. When I tested frontier general-purpose models, my assumption collapsed. They already possessed much of the required capability. What they needed was the right workflow.

Let me qualify that claim by pointing out that one excellent result does not establish reliability across every genre, length, and editorial standard. The American Society for Indexing’s 2026 trajectory report remains skeptical of LLM-generated indexes, which can appear convincing while concealing serious omissions and structural failures.

But I am now convinced that software can play a much larger role in subject indexing than I had previously considered. Modern language models can perform much of its initial intellectual construction. The human no longer has to begin with a blank page.

A generated answer is not a publishing process.

A prompt is not a publishing process

You can give a powerful language model a book and ask for an index. The answer may be impressive. But a single response has not recovered the real structure of the final PDF, kept the whole book in view, reconciled repeated ideas, checked the finished hierarchy, or shown why each locator belongs.

IndexPDF treats those responsibilities as a long-running editorial process. It first prepares the final PDF and recovers its document structure. It constructs index material across the book, reconciles the pieces, audits and edits the complete subject index, grounds retained locators in exact source evidence, and validates the result. Project-specific estimates show that this work continues in the background rather than pretending that a finished index appears immediately.

  1. Prepare the final PDF and recover the book’s readable structure.
  2. Construct and reconcile index material across the publication.
  3. Audit and edit the complete subject index.
  4. Organize headings, subdivisions, terminology, and cross-references.
  5. Ground locators in exact source evidence.
  6. Validate the result and surface unresolved questions for review.

The handoff is therefore not a bag of candidates waiting to be rebuilt. It is a substantially finished index whose structure and evidence have already received book-wide attention. That does not make every editorial question disappear. It means the remaining questions can be reviewed as questions, not excavated from an unexamined draft.

The index is edited before you see it

Subject-index quality depends on relationships that are visible only when the index is read as a whole. A useful heading can be weak because its subheadings repeat one another. Two sensible entries can provide redundant access. A cross-reference can send the reader through an unnecessary detour. Several locally plausible terms can conceal one inconsistent vocabulary.

IndexPDF performs a whole-index editorial pass before handoff. It can consolidate fragmented access without erasing real distinctions, improve heading and subheading organization, reconcile terminology and style, remove unnecessary cross-reference detours, and preserve alternate access that helps a reader. The aim is not merely to accumulate correct-looking entries. It is to edit the subject index as a navigation system.

Synthetic before-and-after subject index showing fragmented Grace and Salvation headings reorganized into a coherent hierarchy with preserved locators and one useful cross-reference.

Synthetic example: whole-index editing can consolidate fragmented headings and remove a redundant cross-reference loop while preserving meaningful distinctions and source locators.

Synthetic example: whole-index editing can consolidate fragmented headings and remove a redundant cross-reference loop while preserving meaningful distinctions and source locators.

Because retained locators remain connected to exact passages, unresolved grounding or organization questions can be surfaced instead of silently accepted. That evidence-grounded relationship is also what makes the later verification workflow inspectable.

Choose the review depth

After the automated editorial work is complete, the reviewer chooses how deeply to inspect it:

IndexPDF review strategy chooser with Full review and Priority review presented as two selectable options.

The actual review-strategy chooser describes scope directly: Full review covers every eligible page, while Priority review focuses on blockers and unresolved subject-index issues without marking pages complete.

The actual review-strategy chooser describes scope directly: Full review covers every eligible page, while Priority review focuses on blockers and unresolved subject-index issues without marking pages complete.
  • Priority review focuses on currently flagged verification blockers and unresolved subject-heading issues. It is an expedited quality-assurance path; it does not mark every page complete.
  • Full review covers every eligible page and every generated mention, including pages where no mention was found. It is exhaustive in workflow scope, not a guarantee that human error or every publication-specific issue is impossible.

A reviewer can switch between Priority and Full review without losing earlier decisions or full-review progress. The choice changes the inspection path, not the underlying index.

The author and the map

This division of labor may also shift who is best positioned to approve an index.

Professional indexers understand access paths, reader behavior, index structure, and the constraints of a publisher’s style. But the author knows the territory. The author knows which distinction is foundational, which phrase is merely incidental, whether two terms name the same idea, and where an argument matters even though it is never announced.

Historically, asking authors to index their own books often meant asking them to learn a second profession at the worst possible moment in publication. When IndexPDF performs substantial construction and whole-index editing first, the author can focus on project-specific judgment: resolving uncertainty, applying the book’s terminology, and approving the result.

The new division of labor is simple:

  • IndexPDF constructs, reconciles, edits, grounds, and validates the index.
  • The workspace keeps the index connected to the passages behind it.
  • The author, editor, or indexer resolves project-specific uncertainty and approves the result.

Professional indexers remain valuable for specialist judgment, house style, difficult projects, and independent quality assurance. Many authors and publishers will prefer to hire one. The change is that the reviewer can begin with an intellectually constructed, book-wide artifact instead of following behind the software to rebuild it.

Export and handoff

The final workspace keeps headings, subheadings, locators, groups, and cross-references together. Cross-reference destinations are navigable, so a reviewer can inspect the route a reader will follow before export.

IndexPDF currently exports the index as Markdown. That portable editorial handoff preserves hierarchy, locators, groups, and cross-references for downstream production. It is not direct Word or InDesign output, typesetting, or an automatic substitute for the publisher’s final style and proofing work. See current pricing and the workflow for publishers.

What “finished” means

A substantially finished index is intellectually constructed rather than delivered as a candidate list. It has been organized across the whole subject index, grounded in the book, validated, and prepared for the reviewer’s chosen quality-assurance path. It can be exported for downstream production without requiring the user to reconstruct its basic architecture.

Finished does not mean automatically approved, guaranteed complete, typeset, or exempt from publisher house style and final proofreading. A polished-looking artifact can still conceal a publication-specific mistake. Someone who understands the book remains responsible for choosing the review depth, resolving uncertainty, and approving the index for publication.

Seven years ago, I tried to grant myself these three wishes through parsers, regular expressions, and a great deal of stubbornness. Two worked. The third required a kind of software that did not yet exist; but now it does. The impossible third index is possible. IndexPDF is built to carry that work substantially farther before asking a human to take responsibility for the final map.

Sources and further reading