Portfolio Document Parsing Data Cleansing The Problem Extract structured data from predictable PDFs and fix extraction errors automatically. The Approach Python parsing + LLM error correction. Minimal infra, co-located with FileMaker. ~150 hours, half on golden datasets. Even "simple" extraction projects require rigorous validation data. The golden dataset makes everything else trustworthy. Data Cleansing 7 / 17