IndexPDF Standards
Subject Index Evaluation
This working draft applies the separately published IndexPDF Subject Indexing Rules 0.2 to a finished back-of-book subject index. The evaluation standards do not define how an index is created.
The default evaluation produces six weighted dimensions and thirteen independent release gates; profiles may add delivery or accessibility claims without changing the default score.
1. Purpose
This standard defines how IndexPDF evaluates and scores a finished back-of-book subject index against its source publication.
An evaluation under this standard answers four questions:
- What subjects and reader tasks does the source require the index to support?
- Are the index's headings and locators supported by the cited pages?
- Can readers find every important treatment that should be indexed?
- Does the index work as a coherent navigation system?
The result is not merely a numerical score. A complete result consists of:
- six dimension ratings and a weighted score out of 100;
- item-level findings and explanations;
- uncertainty and audit-completeness information; and
- independent publication-readiness gates.
This document is normative for an evaluation only when the evaluation report explicitly identifies IndexPDF Subject Index Evaluation Standard 0.2.
2. What this standard covers
The standard applies to subject indexes for books and comparable long-form publications. It evaluates the delivered index, regardless of whether the index was created by a person, software, or a combined process.
The default evaluation covers:
- meaningful subject coverage;
- editorial selectivity;
- conceptual and stance fidelity;
- heading and subheading architecture;
- locator existence, scope, treatment, and complete-path fit;
- missing access and reader tasks;
- cross-references and alternate access;
- density, arrangement, mechanics, and whole-index coherence; and
- the validity and completeness of the evaluation evidence.
The default score does not evaluate the candidate creator's internal workflow, AI policy, commercial process, data governance, or release approvals. Digital-link behavior, EPUB semantics, PDF/UA, WCAG conformance, multilingual collation, and other delivery requirements are evaluated only when a named profile is added to the report. Profile results are reported separately from the default subject-index score.
Name, author, Scripture, geographic, cumulative, database, citation, search-ranking, and journal-inclusion indexes require separate or additional profiles.
3. Governing principles
| Principle | Requirement |
|---|---|
| Source grounded | Every substantive judgment must be based on the exact source edition and the delivered candidate index. |
| Source first | The expected-subject benchmark is created and reviewed without seeing the candidate index. |
| Complete path | A locator is judged against the complete main-heading and subheading path, not an isolated keyword. |
| Two-way audit | The evaluation checks both candidate-to-source support and source-to-candidate missing access. |
| Complete denominators | A full audit accounts for every required locator, expected treatment, reader task, cross-reference, and structure item. |
| Independent diagnostics | Treatment depth and complete-path fit are reported separately, even though the lower value sets the displayed locator grade. |
| Gates remain independent | A high score cannot cancel a critical defect. |
| Explicit uncertainty | Uninspectable evidence is reported as uncertainty, not guessed as a pass or failure. |
| Reproducibility | The report identifies the source, policy, benchmark, audit mode, and calculation profile. |
4. Source and audit scope
Exact source
IPDF-SRC-01The evaluation SHALL identify the exact source edition and files, document-page span, printed or logical page labels, and page map. Roman numerals, prefixes, and irregular labels SHALL be preserved as labels rather than silently converted.
Indexable matter
IPDF-SRC-02Every source region SHALL be classified as included, excluded, conditional, or unavailable before the benchmark is frozen. On a mixed page, eligible and excluded regions SHALL be treated separately.
Substantive body text, quotations, inspectable notes, captions, and table text are included by default. Contents pages, running furniture, bibliographies or source lists, publisher indexes, production matter, and unavailable material are excluded by default unless the evaluation policy states otherwise.
Source quality
IPDF-SRC-03OCR or extraction quality SHALL be checked before automated analysis. Material defects SHALL be corrected or recorded as uninspectable.
Approved chunks
IPDF-SRC-04The source SHALL be divided into approved intellectual units with complete, nonoverlapping judgment ownership. Context may overlap, but each in-scope page has exactly one owner for audit accounting.
The normative candidate-quality, locator, architecture, and presentation requirements are published separately as the IndexPDF Subject Indexing Rules 0.2. An evaluation under this standard SHALL apply that immutable edition and cite applicable rule anchors in its findings.
5. Evaluation sequence
A full evaluation follows this sequence:
- Identify the source and construct the page map.
- Define and approve source chunks.
- Freeze the evaluation policy.
- Discover expected subjects and reader tasks without seeing the candidate.
- Synthesize and independently review the benchmark.
- Freeze the benchmark.
- Mechanically normalize the delivered candidate without editorial repair.
- Account for every delivered path, locator, range, and cross-reference.
- Audit each locator assignment against the source.
- Audit every benchmark treatment for missing access.
- Audit global structure, cross-references, density, mechanics, and coherence.
- Calculate the six dimensions deterministically.
- Produce item explanations, the scorecard, gates, limitations, and the public report.
The benchmark is candidate-blind. Candidate wording or structure may not be used to decide what the source ought to have indexed.
Candidate normalization is mechanical. The evaluator may resolve layout structure, but may not repair, deduplicate, rewrite, or reinterpret the delivered index before scoring it.
6. Locator evaluation
Each assessable locator receives three related but distinct judgments.
6.1 Treatment depth
Treatment depth asks how much independently useful information the destination supplies about the complete heading path.
| Treatment category | Diagnostic value |
|---|---|
| Substantive | 1.00 |
| Mixed | 0.70 |
| Weak or contentless presence | 0.25 |
| Absent or invalid destination | 0.00 |
attribution_only, citation_only, and incidental_example are weak only when the passage supplies no independently useful information about the complete path.
6.2 Complete-path fit
Complete-path fit asks whether the entire heading path accurately describes the treatment.
| Fit category | Diagnostic value |
|---|---|
| Exact | 1.00 |
| Material partial | 0.70 |
| Minor mismatch | 0.35 |
| Major mismatch | 0.15 |
| No fit | 0.00 |
The displayed locator grade is:
This grade explains the locator's strengths and limitations. It is not itself the locator's rating credit.
6.3 Keep decision and rating credit
| Decision | Meaning | Reliability credit |
|---|---|---|
supported | Keep the delivered locator unchanged | 1 |
partially_supported | Relevant evidence exists, but do not keep the locator as delivered | 0 |
unsupported | Do not keep the locator as delivered | 0 |
uninspectable | Evidence cannot be judged reliably | No central credit; neutral 0–1 bound |
not_measured | Required audit work is incomplete | Blocks a full calculation |
A locator may be supported with mixed treatment. For example, a supported locator with mixed treatment and exact fit has a displayed diagnostic grade of 70 and reliability credit of 1. A weak, contentless mention cannot be kept. Strong treatment with a material relationship or stance mismatch cannot be kept as delivered.
7. Page-reference Reliability
This standard uses binary keep precision:
Expected-treatment recall is:
The base Page-reference Reliability rating is:
This balances two risks:
- precision prevents unsupported or materially misdescribed locators from receiving credit; and
- recall prevents a sparse index from scoring highly merely because its few locators are accurate.
Treatment and fit diagnostics remain visible at item level but are not averaged into keep precision.
8. Six scored dimensions
Each dimension is rated on a five-point scale. Dimension points are calculated from the rating and weight, and the six point values sum to a maximum of 100.
| Dimension | Weight | What it measures |
|---|---|---|
| Meaningful Coverage | 20 | Whether essential, major, and eligible optional subjects receive useful access; importance weights are 3, 2, and 1, with complete access 1, partial access 0.5, and missing access 0 |
| Editorial Selectivity | 15 | Whether locators favor meaningful treatment over weak presence, plus whether index density is reasonably distributed across the source |
| Conceptual / Stance Fidelity | 15 | Whether headings preserve the source's concepts, relationships, compound meaning, attribution, and degree of certainty |
| Page-reference Reliability | 25 | The harmonic mean of binary keep precision and expected-treatment recall, subject to independent safeguards |
| Findability / Navigation | 20 | Whether reader tasks succeed, heading architecture supports access, and cross-references work |
| Mechanics / Consistency | 5 | Whether the delivered index is structurally valid, consistent, readable, and mechanically usable |
8.1 Editorial Selectivity
Editorial Selectivity contributes up to 15 points:
- up to 10 points from treatment selectivity; and
- up to 5 points from density fit.
Treatment-selectivity credit is 1 for substantive treatment, 0.5 for mixed treatment, and 0 for weak or contentless presence. The density component uses two chapter-level measures, weighted equally and aggregated by indexable source word count:
| Density measure | Ideal range per 1,000 indexable source words | Acceptable range |
|---|---|---|
| Locator-bearing complete heading paths | 6–10 | 4–12 |
| Expanded locator occurrences | 15–25 | 10–30 |
Density is a calibration measure, not a quota. It does not decide which subjects should be indexed, and a threshold alone does not establish overindexing or underindexing.
8.2 Conceptual Fidelity and Mechanics
Conceptual-fidelity and heading-architecture nodes use these credits: pass 1, minor issues 0.85, major issues 0.55, and fail 0.
Mechanics nodes use: pass 1, cosmetic issues 0.95, minor issues 0.85, major issues 0.55, and fail 0.
8.3 Findability
Findability combines:
- reader-task success: 60%;
- heading and access architecture: 30%; and
- cross-reference validity: 10%.
If cross-references are genuinely inapplicable, the first two components are renormalized to two-thirds and one-third.
8.4 Caps and rounding
Dimension ratings are rounded to the nearest half point after applicable caps and uncertainty rules. Editorial Selectivity uses its documented 10-plus-5 point calculation. Weighted points are rounded to two decimal places.
Caps prevent a strong average from concealing high-value omissions, distributed unsupported patterns, major stance failures, reader-task failures, defective access architecture, or systematic mechanical problems.
Item grades are explanatory. They SHALL NOT be averaged to reconstruct a dimension rating or total score.
9. Critical gates
The score and gates are separate results. Gates do not subtract points; they identify failures too serious to be concealed by an average.
Out-of-scope locator
IPDF-GATE-001Fabricated, nonexistent, or out-of-scope locator
Systematic unsupported locator pattern
IPDF-GATE-002Systematic incidental or unsupported locator pattern
Central omission
IPDF-GATE-003Central subject or conclusion materially omitted
Source-stance distortion
IPDF-GATE-004Heading reverses or seriously misrepresents source stance
Unsupported compound heading
IPDF-GATE-005Compound-heading locators support only separate components
Invalid see substitution
IPDF-GATE-006A see reference replaces a warranted substantive entry
Invalid cross-reference
IPDF-GATE-007Unresolved, self-referential, circular, or chained cross-reference
Excessive heading depth
IPDF-GATE-008Third-level heading under the baseline display profile
Systematic clutter
IPDF-GATE-009Systematic named-entity, example, or citation clutter
Unresolved grounding failure
IPDF-GATE-010Critical or major unresolved grounding failure
Excessive uninspectability
IPDF-GATE-011More than 1% of in-scope locator assignments are uninspectable without a frozen alternative tolerance
Wrong source span
IPDF-GATE-012Wrong source span
Invalid or incomplete structure
IPDF-GATE-013Structurally invalid or incomplete output
A complete public result SHALL report every gate as pass, fail, not applicable, or otherwise explicitly unresolved. A numerical score without gate status is not a complete result under this standard.
10. Audit completeness and uncertainty
Full audit
A full audit accounts for every frozen denominator exactly once. Required locators, treatments, subjects, tasks, references, and structure nodes may not remain not_measured.
Pilot or sampled review
A pilot may be used to calibrate judgments or identify likely problems. It SHALL NOT be represented as a full-index score or full conformance result. Its population, selection method, and unexamined areas must be disclosed.
Uninspectable evidence
uninspectable is used when the destination or evidence cannot be judged reliably. It is not silently counted as success or failure. The calculation reports lower and upper bounds. A score is withheld when uncertainty could change the rounded rating or applicable cap.
11. Evidence and reproducibility
The calculation uses validated structured records. Explanatory prose may clarify a judgment but cannot alter its category or arithmetic.
A reproducible evaluation identifies:
- the exact source and page map;
- the approved source chunks and inclusion policy;
- the frozen evaluation policy;
- the candidate-blind benchmark;
- the normalized candidate and complete item inventory;
- the locator, missing-access, cross-reference, and structure audits;
- the scoring rubric and calculation profile;
- audit mode, uncertainty policy, deviations, and limitations; and
- dimension results, caps, gates, and item explanations.
Hashes join exact records and help detect accidental input mix-ups. They are content identities, not security attestations.
Two evaluations are directly comparable only when their source, benchmark, policy, page map, chunks, inclusion scope, audit mode, uncertainty policy, rubric, and calculation profile are compatible.
12. Public report requirements
A public report under this standard SHALL include:
- candidate and source identification;
- this standard's version;
- scoring-rubric and calculation-profile identities;
- full or pilot audit status;
- all six ratings, weights, and awarded points;
- the total score and uncertainty status;
- all critical gate outcomes;
- raw counts and denominators needed to interpret precision and recall;
- concise item-level explanations and rule identifiers;
- material limitations or deviations; and
- the evaluation date.
Public reports should not reproduce protected source text or expose secrets, private file paths, or storage identifiers. Evidence may be represented by stable internal identifiers and public-safe summaries.
13. How to cite an evaluation
Use this statement for an evaluation that satisfies the requirements above:
This subject index was evaluated under the IndexPDF Subject Index Evaluation Standard 0.2. The evaluation report states its audit mode, source and benchmark identities, score, uncertainty, and critical-gate results.
A shorter website label may be used when it links to the full report and this standard:
Scored under the IndexPDF Subject Index Evaluation Standard 0.2.
14. Authority
The editorial criteria are informed by ANSI/NISO Z39.4-2021, established professional indexing guidance, and applicable technical standards. The normative scoring requirements are the provisions of this document.