Sharing My Handy Tool to Extract Monthly Returns
Automating the Boring Stuff
Link to the tool: PDF → Monthly Returns (Free)
Acceptable PDF format:
Text-based PDFs work best. Scanned image-only PDFs likely to fail unless the text layer exists (OCR).
Unencrypted PDFs (no password).
Monthly returns table provided in PDF in the following format:
User Guide:
Toggle source return format in percentage or aboslute.
Upload one or more PDF factsheets.
The tool extracts monthly, full-year, and calculated yearly returns. It checks full-year value against calculated yearly returns. If the two differ, it will remove the entire year’s returns.
The tool automatically labels columns with the fund name (from the filename, so name your files before uploading).
Outputs appear in tables instantly, with a CSV download file.
Why I Built This Tool
For years I carried the same frustration many allocators know too well: building and maintaining my own return database.
Every month, inboxes fill with PDF factsheets. Each manager has their own template. Some put monthly numbers in a table, some bury them in text, others change layouts mid-year. Extracting the data was a grind.
My workflow used to look like this:
Manual entry into spreadsheets — prone to errors, and always behind schedule.
Copy/paste gymnastics with PDF readers that mangled columns.
Custom Excel macros that broke whenever formatting shifted.
The result: analysis delayed, energy drained, and less time for what actually matters — making allocation decisions.
How This Helps Allocators
Faster database upkeep: One upload and you have structured data.
Cleaner analysis: Numbers arrive in consistent tables you can merge.
Error reduction: No more typos from manual entry.
Current Shortcomings
Layout variability: Some PDFs use exotic formatting that confuses the parser. You may still see occasional missing or misaligned numbers.
No universal standard: Every manager’s factsheet looks different. The tool handles most, but edge cases will always exist.
Fund name detection: Right now column headers come from the file name. That works, but it’s not as elegant as parsing the canonical fund name from the PDF itself.
Performance: For very large uploads or unusual PDFs, processing may take longer than expected.
Privacy & Data Protection
No storage of PDFs: Uploaded factsheets are processed in-memory for return extraction. Copies of the original files are not stored or retained on my server.
No distribution of fund documents: I do not retain, distribute, or share the original fund tear sheets. Your fund managers’ materials remain private to you.
Transparency: I believe in being clear about what happens behind the scenes. If I change how data is handled, I will tell you.


