4. Results

4.1 Pooled results

Table 6. Pooled extraction results.
Measure Pooled result
Original PDF records 2,484
Evaluable matched records 2,482
Pipeline success 99.92%
Correct atomic facts 158,715
Partially correct facts 13,961
Unsupported/contradicted facts 2,158
Missing facts 80,371
Unverifiable facts 1,705
Strict precision / recall / F1 90.8% / 62.7% / 74.2%
Weighted precision / recall / F1 94.8% / 65.5% / 77.4%
Completeness 68.2%
Unsupported-content rate 1.23%
Critical unsupported-claim rate 2.78%
Schema completeness 72.4%

The principal result is a 29.3-percentage-point gap between weighted precision and weighted recall. The extraction is conservative: facts that appear in the output are usually supported, but a substantial fraction of source information is omitted. The pooled unsupported-content rate is low, yet low hallucination does not compensate for missing retrieval evidence when the system is used to search for candidates.

Under the common thresholds, 201 documents are successful, 1,989 are partially successful, and 294 fail. The pooled pipeline-success rate is 99.92%, not exactly 100%, because Banking and Business Development each report one unmatched or non-evaluable record. The outcome denominator is the full set of 2,484 source PDFs, whereas factual metrics are computed from the available atomic evidence.

Figure 2. Pooled document outcomes.
Figure 2. Pooled document outcomes.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Most documents are partially successful. The small fully successful fraction is caused by the joint threshold: a document must simultaneously achieve high F1, high completeness, low unsupported content, and no critical unsupported or contradicted claim.

Limitation. Outcome labels depend on the operational thresholds in prompt2(5); a different application could choose different cutoffs.

Open SVG

4.2 Dataset-level variation

Table 7. Cross-dataset distribution of weighted F1.
Statistic Across 24 categories
Macro mean weighted F1 78.3%
Median weighted F1 82.3%
Standard deviation 0.127
First quartile 76.5%
Third quartile 85.9%
Minimum / maximum 43.5% / 90.9%
Macro mean weighted precision 95.2%
Macro mean weighted recall 67.9%
Table 8. Highest- and lowest-performing datasets.
Group Dataset CVs W. precision W. recall W. F1 Completeness Failed
Highest TEACHER 102 97.8% 84.9% 90.9% 85.6% 1.0%
Highest AUTOMOBILE 36 100.0% 81.6% 89.8% 81.6% 0.0%
Highest BUSINESS DEVELOPMENT 120 98.5% 82.5% 89.8% 83.7% 1.7%
Highest SALES 116 99.8% 79.0% 88.2% 79.0% 1.7%
Highest AGRICULTURE 63 93.3% 81.9% 87.2% 87.7% 0.0%
Lowest ACCOUNTANT 118 89.5% 28.7% 43.5% 32.1% 89.8%
Lowest BPO 22 97.6% 31.9% 48.1% 32.2% 81.8%
Lowest INFORMATION TECHNOLOGY 120 75.9% 38.6% 51.1% 50.8% 46.7%
Lowest HR 110 89.4% 57.8% 70.2% 64.5% 16.4%
Lowest CHEF 118 96.8% 61.8% 75.5% 62.8% 9.3%

Teacher, Automobile, Business Development, Sales, and Agriculture form the highest-performing group by weighted F1. Accountant, BPO, Information Technology, HR, and Chef form the lowest-performing group. The ordering is driven primarily by recall and completeness rather than unsupported content. This pattern implies that layout, vocabulary, section conventions, and prompt-to-document alignment vary materially by occupation.

Figure 3. Weighted F1 by occupational dataset.
Figure 3. Weighted F1 by occupational dataset.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Weighted F1 varies widely across occupational datasets. A single pooled score therefore cannot be assumed to generalize uniformly to every CV domain.

Limitation. Dataset sizes differ, and weighted F1 alone does not identify whether a low score is caused by omission, unsupported content, or both.

Open SVG
Figure 4. Dataset precision-recall profile.
Figure 4. Dataset precision-recall profile.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Most datasets occupy a high-precision but lower-recall region. Accountant, BPO, and Information Technology are prominent low-recall outliers, supporting the interpretation of conservative extraction with substantial omission.

Limitation. Each point aggregates heterogeneous documents; confidence intervals and language-stratified results are unavailable.

Open SVG
Figure 5. Completeness versus weighted F1.
Figure 5. Completeness versus weighted F1.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Completeness and weighted F1 are strongly associated across the 24 categories (Pearson r = 0.958). Missing source facts are therefore a principal determinant of overall quality.

Limitation. Correlation does not establish that completeness alone causes F1; both measures share atomic-fact counts and are mathematically related.

Open SVG
Figure 6. Schema completeness versus weighted F1.
Figure 6. Schema completeness versus weighted F1.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Schema completeness has little category-level association with weighted F1 (r = 0.080). A populated JSON structure is not evidence that its values are correct or complete.

Limitation. Schema fields and factual facts are not equally numerous across datasets; the correlation is descriptive.

Open SVG
Figure 7. Critical unsupported-claim rate by dataset.
Figure 7. Critical unsupported-claim rate by dataset.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Critical unsupported claims are concentrated in a small number of datasets. Teacher has a comparatively high critical-document rate despite strong F1, demonstrating why critical-error checks must remain separate from average factual metrics.

Limitation. Critical status depends on the error taxonomy and should be independently audited for high-stakes deployment.

Open SVG
Figure 8. Outcome composition by dataset.
Figure 8. Outcome composition by dataset.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Dataset outcome composition reveals that some datasets contain no fully successful documents even when most documents are partially usable. Accountant and BPO have particularly high failed shares.

Limitation. The outcome categories combine technical and content conditions and are threshold-sensitive.

Open SVG

4.3 Segment-level performance

Table 9. Segment-level extraction performance.
Segment Applicable W. precision W. recall W. F1 Completeness Unsupported Schema
Experience 2,482 95.7% 81.8% 88.2% 85.3% 0.25% 92.3%
Profile 2,444 94.7% 75.2% 83.8% 79.1% 0.38% 89.0%
Skills 2,473 97.0% 64.1% 77.2% 66.0% 0.18% 88.1%
Education 2,462 93.3% 56.7% 70.5% 60.5% 0.36% 86.6%
Certifications 1,147 90.9% 51.1% 65.5% 55.9% 0.56% 71.6%
Languages 451 90.8% 45.8% 60.9% 50.3% 0.34% 80.3%
Document metadata 1,976 49.7% 18.7% 27.2% 19.7% 47.64% 55.9%
Category classification 2,329 79.9% 7.1% 13.1% 8.9% 0.18% 27.0%
Normalized fields 2,118 69.2% 2.9% 5.5% 3.0% 27.52% 5.4%
Seniority 2,252 53.9% 2.4% 4.6% 3.1% 30.46% 22.1%

Experience is the strongest major segment, with weighted F1 of 88.2% and completeness of 85.3%. Profile and Skills are also useful, although Skills recall remains substantially below precision. Education, Certifications, and Languages show progressively lower recall. Category classification, normalized fields, seniority, and document metadata perform too poorly to be treated as authoritative fields.

Document metadata is unusual because it combines low recall with a very high unsupported-content rate. Normalized fields and seniority also contain material unsupported content. These fields should either be removed from hard filters, recomputed deterministically from source evidence, or displayed only after verification.

Figure 9. Segment-level weighted precision, recall, and F1.
Figure 9. Segment-level weighted precision, recall, and F1.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. The core narrative segments have high precision but varying recall. Derived and classification-oriented segments show extremely low recall and F1.

Limitation. Segment definitions aggregate multiple fields; performance of an individual field may differ from the segment average.

Open SVG
Figure 10. Segment-level fact composition.
Figure 10. Segment-level fact composition.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Omitted facts dominate weak segments. Unsupported facts are less frequent overall but are concentrated in document metadata, normalized fields, and seniority.

Limitation. Counts are affected by the number of applicable facts in each segment and should not be interpreted as equal-risk denominators.

Open SVG
Figure 14. Dataset-by-segment weighted-F1 heatmap.
Figure 11. Dataset-by-segment weighted-F1 heatmap.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. The heatmap shows occupation-specific segment failures rather than a uniform global weakness. This supports dataset-aware monitoring and targeted prompt or parser remediation.

Limitation. Cells summarize segment aggregates and do not display document-level uncertainty.

Open SVG

4.4 Error structure and severity

Table 10. Most frequent error categories.
Error category Count Share of logged errors
Omitted fact 72,294 68.02%
Empty field 11,622 10.94%
Partially correct fact 10,577 9.95%
Segmentation error 2,643 2.49%
Normalization error 1,719 1.62%
Evaluation uncertainty 1,471 1.38%
Missing segment 1,420 1.34%
Unsupported fact 945 0.89%
Incorrect seniority 622 0.59%
Under-specification 581 0.55%
Cross-field inconsistency 554 0.52%
Contradicted fact 457 0.43%
Incorrect category 318 0.30%
Incorrect language 273 0.26%
Over-generalization 187 0.18%

Omitted fact is the dominant logged error (72,294 occurrences), followed by empty field and partially correct fact. Segmentation errors, normalization errors, evaluation uncertainty, missing segments, unsupported facts, and incorrect seniority form a second tier. The error profile is therefore structurally consistent with the aggregate precision-recall gap: the system is more likely to omit or under-specify than to invent.

Figure 11. Error Pareto distribution.
Figure 12. Error Pareto distribution.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. A small number of error categories account for most logged errors. Remediation should prioritize omitted facts, empty fields, partial facts, and segmentation before less frequent normalization and classification errors.

Limitation. Frequency is not severity. A rare critical identity error may be more consequential than many minor omissions.

Open SVG
Figure 16. Error severity distribution.
Figure 13. Error severity distribution.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Most logged errors are moderate or major; critical errors are rare in count but cannot be ignored because they can alter identity, employment, education, certification, or date evidence.

Limitation. Severity labels depend on the evaluation taxonomy and may require application-specific recalibration.

Open SVG

4.5 Document-level variation

Dataset averages conceal considerable document-level heterogeneity. Some CVs achieve high F1 and completeness in otherwise weak datasets, while some fail within strong datasets. Document layout, OCR quality, non-standard headings, multilingual text, tables, columns, and long employment histories can all affect extraction.

Figure 12. Distribution of document-level weighted F1.
Figure 14. Distribution of document-level weighted F1.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. The document-level F1 distribution is broad and is not well represented by a single mean. The mass below the success threshold explains the low fully successful share.

Limitation. The distribution does not identify the document characteristics responsible for each score.

Open SVG
Figure 13. Document-level completeness versus weighted F1.
Figure 15. Document-level completeness versus weighted F1.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Documents with higher completeness generally have higher weighted F1, but the scatter shows that factual quality also depends on partial and unsupported facts.

Limitation. The axes share component counts; visual association should not be treated as an independent causal estimate.

Open SVG
Figure 15. Multimetric dataset quality heatmap.
Figure 16. Multimetric dataset quality heatmap.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. The multimetric heatmap highlights the difference between precision, recall, completeness, schema population, failure rate, and critical risk. No single metric is sufficient for deployment acceptance.

Limitation. Color scaling can visually amplify small differences; numerical values should be consulted for decisions.

Open SVG
Figure 17. Dataset volume versus weighted F1.
Figure 17. Dataset volume versus weighted F1.

Source: pooled PDF-grounded evaluation workbooks; calculations in the combined analysis.

Interpretation. Category size has only a weak association with weighted F1 (r = 0.118). Poor performance is not explained simply by having more CVs.

Limitation. This is an ecological dataset-level correlation and does not test document length or token count.

Open SVG