Common PDF-grounded audit
Four extractors under identical measurement rules
| System | Outputs | Weighted precision | Weighted recall | Weighted F1 | 95% bootstrap CI | Completeness | Unsupported |
|---|---|---|---|---|---|---|---|
| Gemma | 2,482/2,484 | 99.40% | 81.16% | 89.36% | 88.93%–89.78% | 81.57% | 0.10% |
| Qwen | 2,483/2,484 | 99.04% | 83.02% | 90.32% | 89.93%–90.73% | 83.72% | 0.13% |
| Hybrid NLP | 2,483/2,484 | 99.25% | 75.78% | 85.94% | 85.27%–86.56% | 76.25% | 0.13% |
| BERT MULTI | 2,483/2,484 | 99.44% | 55.83% | 71.51% | 70.63%–72.38% | 56.09% | 0.09% |
The four values are evaluator-compatible. They remain conditional on automated PDF text extraction, lexical support thresholds, fact splitting and ChatGPT-assisted omission-candidate generation.
Category-stratified comparison
All 96 system–category cells and all six contrasts
| Contrast | Mean ΔF1 | Median ΔF1 | Category wins |
|---|---|---|---|
| Gemma − Qwen | -0.90 pp | -0.94 pp | 4–20; ties 0 |
| Gemma − Hybrid NLP | +3.83 pp | +4.23 pp | 21–3; ties 0 |
| Gemma − BERT MULTI | +17.93 pp | +17.54 pp | 24–0; ties 0 |
| Qwen − Hybrid NLP | +4.73 pp | +5.33 pp | 22–2; ties 0 |
| Qwen − BERT MULTI | +18.83 pp | +18.84 pp | 24–0; ties 0 |
| Hybrid NLP − BERT MULTI | +14.10 pp | +13.37 pp | 24–0; ties 0 |
Field specialization
Segment behavior is compared jointly
A global score can conceal field-specific strengths, omissions and evaluator sensitivity. Segment patterns should guide routing hypotheses, then be verified by human adjudication.
Interpretation
Results describe complete pipelines
HNLP’s result includes PDF-to-Markdown structure, schema aliases, dictionaries, synonyms, fuzzy boundaries, evidence dominance, normalization, and validation. BERT’s result includes model identity, tokenizer, chunking, labels, threshold, fallback, sanitization, duplicate resolution, mapping, grouping, splitting, backfill, and validation. The benchmark does not isolate these components.