Open source
bloom-parser
A service that turns images, PDFs and spreadsheets into one structured document.
I built it and published it as open source.
- Role
- Sole author
- Context
- Personal project, open source
- When
- 2026
What I did
- Routed 11 input formats through 4 adapters into one document model.
- Made OCR optional at build time: the default binary needs no native libraries, and a build tag adds Tesseract.
Impact
Public MIT-licensed code; a new format is an adapter plus a detection rule, and the pipeline, OCR and exporters don't change.
When things go wrong
When the OCR engine isn't installed,
OCR requests get a clear per-page error and every other path keeps working.
Built with
Go, gRPC, Python, Tesseract, Vue