Benchmark publication

Critical comparative evaluation of five CV extraction pipelines

ChatGPT, Gemma, Qwen, proprietary Hybrid NLP, and proprietary BERT/GLiNER extraction evaluated with explicit evidence tiers, mathematical models, and source-level limitations.

2,484source CVs

across 24 occupational categories

5pipelines

three generative and two non-generative

90.3%highest conditional F1

Qwen under the common automated evaluator

96system–category cells

24 categories × four extractors under identical rules

Abstract

A critical comparison of five extraction pipelines

This study compares ChatGPT, Gemma, Qwen, Hybrid NLP (HNLP), and NLP BERT MULTI across a 24-category CV corpus. In addition to preserving the historical evidence tiers, it re-evaluates the raw Gemma, Qwen, HNLP and BERT JSON archives against all 2,484 source PDFs using one frozen fact splitter, matcher, threshold set, missing-output policy and aggregation rule.

Interpretive boundary.

ChatGPT proposes omission candidates, but every candidate must be supported by the PDF; it is not treated as truth. The common evaluator permits controlled conditional comparison of the four tested extractors, but no blinded human gold standard establishes intrinsic or causal model superiority.

Evidence architecture

Comparability is treated as a design condition

Figure 1. Five-system evidence architectureThe figure separates valid paired comparisons from descriptive cross-tier profiles.
Figure 2. Historical coverage and common-evaluator F1The final panel places Gemma, Qwen, HNLP and BERT under the same PDF-grounded measurement rules.

Principal findings

The pooled result hides field specialization

Controlled comparison

Qwen has the highest conditional weighted F1 (90.3%) in the common automated audit; the claim is explicitly evaluator-dependent.

All systems together

Every category and segment figure now includes Gemma, Qwen, HNLP and BERT together. Pairwise panels were replaced by four-column heatmaps and deviations from category means.

Academic caution

PDF text extraction, atomic splitting, lexical thresholds and ChatGPT-assisted candidate generation remain possible error sources. A blinded dual-rater subset is required to test rank robustness.

Publication files

Download the complete analysis

NextMethods and evidence tiers →
Figure