across 24 occupational categories
Benchmark publication
Critical comparative evaluation of five CV extraction pipelines
ChatGPT, Gemma, Qwen, proprietary Hybrid NLP, and proprietary BERT/GLiNER extraction evaluated with explicit evidence tiers, mathematical models, and source-level limitations.
three generative and two non-generative
Qwen under the common automated evaluator
24 categories × four extractors under identical rules
Abstract
A critical comparison of five extraction pipelines
This study compares ChatGPT, Gemma, Qwen, Hybrid NLP (HNLP), and NLP BERT MULTI across a 24-category CV corpus. In addition to preserving the historical evidence tiers, it re-evaluates the raw Gemma, Qwen, HNLP and BERT JSON archives against all 2,484 source PDFs using one frozen fact splitter, matcher, threshold set, missing-output policy and aggregation rule.
ChatGPT proposes omission candidates, but every candidate must be supported by the PDF; it is not treated as truth. The common evaluator permits controlled conditional comparison of the four tested extractors, but no blinded human gold standard establishes intrinsic or causal model superiority.
Evidence architecture
Comparability is treated as a design condition
Principal findings
The pooled result hides field specialization
Controlled comparison
Qwen has the highest conditional weighted F1 (90.3%) in the common automated audit; the claim is explicitly evaluator-dependent.
All systems together
Every category and segment figure now includes Gemma, Qwen, HNLP and BERT together. Pairwise panels were replaced by four-column heatmaps and deviations from category means.
Academic caution
PDF text extraction, atomic splitting, lexical thresholds and ChatGPT-assisted candidate generation remain possible error sources. A blinded dual-rater subset is required to test rank robustness.
Publication files