Empirical findings

Four extractors under one common PDF-grounded audit

The results page presents global, category, pairwise and segment comparisons for Gemma, Qwen, HNLP and BERT under identical automated measurement rules.

Common PDF-grounded audit

Four extractors under identical measurement rules

SystemOutputsWeighted precisionWeighted recallWeighted F195% bootstrap CICompletenessUnsupported
Gemma2,482/2,48499.40%81.16%89.36%88.93%–89.78%81.57%0.10%
Qwen2,483/2,48499.04%83.02%90.32%89.93%–90.73%83.72%0.13%
Hybrid NLP2,483/2,48499.25%75.78%85.94%85.27%–86.56%76.25%0.13%
BERT MULTI2,483/2,48499.44%55.83%71.51%70.63%–72.38%56.09%0.09%

The four values are evaluator-compatible. They remain conditional on automated PDF text extraction, lexical support thresholds, fact splitting and ChatGPT-assisted omission-candidate generation.

Category-stratified comparison

All 96 system–category cells and all six contrasts

ContrastMean ΔF1Median ΔF1Category wins
Gemma − Qwen-0.90 pp-0.94 pp4–20; ties 0
Gemma − Hybrid NLP+3.83 pp+4.23 pp21–3; ties 0
Gemma − BERT MULTI+17.93 pp+17.54 pp24–0; ties 0
Qwen − Hybrid NLP+4.73 pp+5.33 pp22–2; ties 0
Qwen − BERT MULTI+18.83 pp+18.84 pp24–0; ties 0
Hybrid NLP − BERT MULTI+14.10 pp+13.37 pp24–0; ties 0
Figure 5. Four-system category weighted F1All four systems are shown together under one common evaluator.
Figure 6. Four-system deviations from category meansEach cell is a system’s F1 minus the four-system mean for that category, in percentage points.

Field specialization

Segment behavior is compared jointly

Figure 7. Four-system segment weighted F1Gemma, Qwen, HNLP and BERT are represented for every canonical segment under the common evaluator.
Pooled-score warning.

A global score can conceal field-specific strengths, omissions and evaluator sensitivity. Segment patterns should guide routing hypotheses, then be verified by human adjudication.

Interpretation

Results describe complete pipelines

HNLP’s result includes PDF-to-Markdown structure, schema aliases, dictionaries, synonyms, fuzzy boundaries, evidence dominance, normalization, and validation. BERT’s result includes model identity, tokenizer, chunking, labels, threshold, fallback, sanitization, duplicate resolution, mapping, grouping, splitting, backfill, and validation. The benchmark does not isolate these components.

NextCritical analysis and trade-offs →
Figure