PDF imports preserve readable text
FPGAi now keeps useful text from PDFs even when a document also contains diagrams, scanned pages, or image-only content.
PDF understanding
- PDFs with selectable text and visual pages now keep both the recovered text and vision-based descriptions, so diagrams do not replace searchable content.
- Scanned PDFs can retain OCR-recovered text alongside visual-page captions, while genuinely image-only PDFs remain available through vision indexing.
Recovery
- If the first PDF enrichment pass times out or returns incomplete output, FPGAi makes a bounded text-recovery attempt before deciding how to index the document.
- The ingester reports recovered page counts and parser details when recovery is needed.