4. Results
4.1 Pooled results
| Measure | Pooled result |
|---|---|
| Original PDF records | 2,484 |
| Evaluable matched records | 2,482 |
| Pipeline success | 99.92% |
| Correct atomic facts | 158,715 |
| Partially correct facts | 13,961 |
| Unsupported/contradicted facts | 2,158 |
| Missing facts | 80,371 |
| Unverifiable facts | 1,705 |
| Strict precision / recall / F1 | 90.8% / 62.7% / 74.2% |
| Weighted precision / recall / F1 | 94.8% / 65.5% / 77.4% |
| Completeness | 68.2% |
| Unsupported-content rate | 1.23% |
| Critical unsupported-claim rate | 2.78% |
| Schema completeness | 72.4% |
The principal result is a 29.3-percentage-point gap between weighted precision and weighted recall. The extraction is conservative: facts that appear in the output are usually supported, but a substantial fraction of source information is omitted. The pooled unsupported-content rate is low, yet low hallucination does not compensate for missing retrieval evidence when the system is used to search for candidates.
Under the common thresholds, 201 documents are successful, 1,989 are partially successful, and 294 fail. The pooled pipeline-success rate is 99.92%, not exactly 100%, because Banking and Business Development each report one unmatched or non-evaluable record. The outcome denominator is the full set of 2,484 source PDFs, whereas factual metrics are computed from the available atomic evidence.
Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Most documents are partially successful. The small fully successful fraction is caused by the joint threshold: a document must simultaneously achieve high F1, high completeness, low unsupported content, and no critical unsupported or contradicted claim.
Limitation. Outcome labels depend on the operational thresholds in prompt2(5); a different application could choose different cutoffs.
Open SVG4.2 Dataset-level variation
| Statistic | Across 24 categories |
|---|---|
| Macro mean weighted F1 | 78.3% |
| Median weighted F1 | 82.3% |
| Standard deviation | 0.127 |
| First quartile | 76.5% |
| Third quartile | 85.9% |
| Minimum / maximum | 43.5% / 90.9% |
| Macro mean weighted precision | 95.2% |
| Macro mean weighted recall | 67.9% |
| Group | Dataset | CVs | W. precision | W. recall | W. F1 | Completeness | Failed |
|---|---|---|---|---|---|---|---|
| Highest | TEACHER | 102 | 97.8% | 84.9% | 90.9% | 85.6% | 1.0% |
| Highest | AUTOMOBILE | 36 | 100.0% | 81.6% | 89.8% | 81.6% | 0.0% |
| Highest | BUSINESS DEVELOPMENT | 120 | 98.5% | 82.5% | 89.8% | 83.7% | 1.7% |
| Highest | SALES | 116 | 99.8% | 79.0% | 88.2% | 79.0% | 1.7% |
| Highest | AGRICULTURE | 63 | 93.3% | 81.9% | 87.2% | 87.7% | 0.0% |
| Lowest | ACCOUNTANT | 118 | 89.5% | 28.7% | 43.5% | 32.1% | 89.8% |
| Lowest | BPO | 22 | 97.6% | 31.9% | 48.1% | 32.2% | 81.8% |
| Lowest | INFORMATION TECHNOLOGY | 120 | 75.9% | 38.6% | 51.1% | 50.8% | 46.7% |
| Lowest | HR | 110 | 89.4% | 57.8% | 70.2% | 64.5% | 16.4% |
| Lowest | CHEF | 118 | 96.8% | 61.8% | 75.5% | 62.8% | 9.3% |
Teacher, Automobile, Business Development, Sales, and Agriculture form the highest-performing group by weighted F1. Accountant, BPO, Information Technology, HR, and Chef form the lowest-performing group. The ordering is driven primarily by recall and completeness rather than unsupported content. This pattern implies that layout, vocabulary, section conventions, and prompt-to-document alignment vary materially by occupation.
Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Weighted F1 varies widely across occupational datasets. A single pooled score therefore cannot be assumed to generalize uniformly to every CV domain.
Limitation. Dataset sizes differ, and weighted F1 alone does not identify whether a low score is caused by omission, unsupported content, or both.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Most datasets occupy a high-precision but lower-recall region. Accountant, BPO, and Information Technology are prominent low-recall outliers, supporting the interpretation of conservative extraction with substantial omission.
Limitation. Each point aggregates heterogeneous documents; confidence intervals and language-stratified results are unavailable.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Completeness and weighted F1 are strongly associated across the 24 categories (Pearson r = 0.958). Missing source facts are therefore a principal determinant of overall quality.
Limitation. Correlation does not establish that completeness alone causes F1; both measures share atomic-fact counts and are mathematically related.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Schema completeness has little category-level association with weighted F1 (r = 0.080). A populated JSON structure is not evidence that its values are correct or complete.
Limitation. Schema fields and factual facts are not equally numerous across datasets; the correlation is descriptive.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Critical unsupported claims are concentrated in a small number of datasets. Teacher has a comparatively high critical-document rate despite strong F1, demonstrating why critical-error checks must remain separate from average factual metrics.
Limitation. Critical status depends on the error taxonomy and should be independently audited for high-stakes deployment.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Dataset outcome composition reveals that some datasets contain no fully successful documents even when most documents are partially usable. Accountant and BPO have particularly high failed shares.
Limitation. The outcome categories combine technical and content conditions and are threshold-sensitive.
Open SVG4.3 Segment-level performance
| Segment | Applicable | W. precision | W. recall | W. F1 | Completeness | Unsupported | Schema |
|---|---|---|---|---|---|---|---|
| Experience | 2,482 | 95.7% | 81.8% | 88.2% | 85.3% | 0.25% | 92.3% |
| Profile | 2,444 | 94.7% | 75.2% | 83.8% | 79.1% | 0.38% | 89.0% |
| Skills | 2,473 | 97.0% | 64.1% | 77.2% | 66.0% | 0.18% | 88.1% |
| Education | 2,462 | 93.3% | 56.7% | 70.5% | 60.5% | 0.36% | 86.6% |
| Certifications | 1,147 | 90.9% | 51.1% | 65.5% | 55.9% | 0.56% | 71.6% |
| Languages | 451 | 90.8% | 45.8% | 60.9% | 50.3% | 0.34% | 80.3% |
| Document metadata | 1,976 | 49.7% | 18.7% | 27.2% | 19.7% | 47.64% | 55.9% |
| Category classification | 2,329 | 79.9% | 7.1% | 13.1% | 8.9% | 0.18% | 27.0% |
| Normalized fields | 2,118 | 69.2% | 2.9% | 5.5% | 3.0% | 27.52% | 5.4% |
| Seniority | 2,252 | 53.9% | 2.4% | 4.6% | 3.1% | 30.46% | 22.1% |
Experience is the strongest major segment, with weighted F1 of 88.2% and completeness of 85.3%. Profile and Skills are also useful, although Skills recall remains substantially below precision. Education, Certifications, and Languages show progressively lower recall. Category classification, normalized fields, seniority, and document metadata perform too poorly to be treated as authoritative fields.
Document metadata is unusual because it combines low recall with a very high unsupported-content rate. Normalized fields and seniority also contain material unsupported content. These fields should either be removed from hard filters, recomputed deterministically from source evidence, or displayed only after verification.
Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. The core narrative segments have high precision but varying recall. Derived and classification-oriented segments show extremely low recall and F1.
Limitation. Segment definitions aggregate multiple fields; performance of an individual field may differ from the segment average.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Omitted facts dominate weak segments. Unsupported facts are less frequent overall but are concentrated in document metadata, normalized fields, and seniority.
Limitation. Counts are affected by the number of applicable facts in each segment and should not be interpreted as equal-risk denominators.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. The heatmap shows occupation-specific segment failures rather than a uniform global weakness. This supports dataset-aware monitoring and targeted prompt or parser remediation.
Limitation. Cells summarize segment aggregates and do not display document-level uncertainty.
Open SVG4.4 Error structure and severity
| Error category | Count | Share of logged errors |
|---|---|---|
| Omitted fact | 72,294 | 68.02% |
| Empty field | 11,622 | 10.94% |
| Partially correct fact | 10,577 | 9.95% |
| Segmentation error | 2,643 | 2.49% |
| Normalization error | 1,719 | 1.62% |
| Evaluation uncertainty | 1,471 | 1.38% |
| Missing segment | 1,420 | 1.34% |
| Unsupported fact | 945 | 0.89% |
| Incorrect seniority | 622 | 0.59% |
| Under-specification | 581 | 0.55% |
| Cross-field inconsistency | 554 | 0.52% |
| Contradicted fact | 457 | 0.43% |
| Incorrect category | 318 | 0.30% |
| Incorrect language | 273 | 0.26% |
| Over-generalization | 187 | 0.18% |
Omitted fact is the dominant logged error (72,294 occurrences), followed by empty field and partially correct fact. Segmentation errors, normalization errors, evaluation uncertainty, missing segments, unsupported facts, and incorrect seniority form a second tier. The error profile is therefore structurally consistent with the aggregate precision-recall gap: the system is more likely to omit or under-specify than to invent.
Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. A small number of error categories account for most logged errors. Remediation should prioritize omitted facts, empty fields, partial facts, and segmentation before less frequent normalization and classification errors.
Limitation. Frequency is not severity. A rare critical identity error may be more consequential than many minor omissions.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Most logged errors are moderate or major; critical errors are rare in count but cannot be ignored because they can alter identity, employment, education, certification, or date evidence.
Limitation. Severity labels depend on the evaluation taxonomy and may require application-specific recalibration.
Open SVG4.5 Document-level variation
Dataset averages conceal considerable document-level heterogeneity. Some CVs achieve high F1 and completeness in otherwise weak datasets, while some fail within strong datasets. Document layout, OCR quality, non-standard headings, multilingual text, tables, columns, and long employment histories can all affect extraction.
Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. The document-level F1 distribution is broad and is not well represented by a single mean. The mass below the success threshold explains the low fully successful share.
Limitation. The distribution does not identify the document characteristics responsible for each score.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Documents with higher completeness generally have higher weighted F1, but the scatter shows that factual quality also depends on partial and unsupported facts.
Limitation. The axes share component counts; visual association should not be treated as an independent causal estimate.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. The multimetric heatmap highlights the difference between precision, recall, completeness, schema population, failure rate, and critical risk. No single metric is sufficient for deployment acceptance.
Limitation. Color scaling can visually amplify small differences; numerical values should be consulted for decisions.
Open SVGSource: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.
Interpretation. Category size has only a weak association with weighted F1 (r = 0.118). Poor performance is not explained simply by having more CVs.
Limitation. This is an ecological dataset-level correlation and does not test document length or token count.
Open SVG